📂 LLM & Models
Comparisons, benchmarks and practical guides on large language models: GPT, Claude, Gemini, Llama and free alternatives.
Xiaomi MiMo-V2.6: Live-streamed RL training that propels an open-source model to #1 on Artificial Analysis
Discover Xiaomi MiMo-V2.6: live RL training, open-source models up to 1T parameters, #1 on Artificial Analysis. Weights available on Hugging Face.
Anthropic accuses GLM-5.3 of "Mythos-class" hacking capabilities — and reignites the war against open weights
Anthropic accuses GLM-5.3 of Mythos-class hacking capabilities and reignites the war against open weights models. Analysis of the explosive report.
MiniMax M3.1 Flash Preview: China continues its war of small, fast models
Discover MiniMax M3.1 Flash Preview, the new fast small Chinese model quietly launched on September 27, 2026, without a model card or press release.
Claude Sonnet 5.5: Anthropic targets the mid-market with a model that's 30% faster and 30% cheaper
Claude Sonnet 5.5: Anthropic targets the mid-market with a model 30% faster and 30% cheaper. An analysis of the new model in the Claude 5.5 family.
Naive-N0.5-Flash: the open-weight 309B MoE with no full-attention layers at all aims for 2,000 tokens/s per user
Naive-N0.5-Flash: NaiveAI's open-weight 309B-parameter MoE without full-attention aims for 2,000 tokens/s per user on Hugging Face.
Gemini 3.8 Live with Live Avatar: Google adds real-time visual presence to its conversational assistant
Google announces Gemini 3.8 Live with Live Avatar: a conversational assistant with real-time visual presence that listens, sees and speaks.
The price war: GPT-6 Sol and Luna at half price, dropped 90 minutes after Claude Opus 5.5
AI price war: OpenAI launches GPT-6 Sol and Luna at half price, just 90 minutes after Claude Opus 5.5's release. Full analysis.
Gemini 3.8 Live: Google's real-time voice at $1.38 per hour — background reasoning and tool execution without interrupting the conversation
Gemini 3.8 Live: Google's real-time AI voice at $1.38/hr. Background reasoning and tool execution without interrupting the conversation.
DeepSeek V4.1-Flash: MIT, 552B, a quarter of the KV-cache, and the end of the Pro model behind the deepseek-flash endpoint as of September 14
DeepSeek V4.1-Flash: MIT weights on Hugging Face, 552B parameters, KV-cache cut to a quarter, and the Pro model ending as of September 14. Full analysis.
OpenAI GPT-6 Astra: 64.6% on Terminal-Bench-Science and ARC-AGI-3 nearly complete, but its written reasoning is becoming harder to monitor
OpenAI GPT-6 Astra scores 64.6% on Terminal-Bench-Science and nears ARC-AGI-3, but its written reasoning becomes harder to monitor.
DeepSeek swaps V4-Pro for V4.1-Flash behind the same endpoint: why you should re-test your pipelines before September 14
DeepSeek replaces V4-Pro with V4.1-Flash behind the same endpoint: find out why you should re-test your pipelines before September 14.
Best Local LLMs (September 2026)
Discover the best local LLMs of September 2026: complete ranking to run AI locally with privacy and no subscription.
# Best Free LLMs (September 2026) Could you also share the body text you'd like translated? I'd be happy to translate it while keeping URLs, /article/xxx slugs, and tool names intact.
Discover the best free LLMs in September 2026: Gemini, DeepSeek, OpenRouter, and more. An honest, up-to-date ranking to avoid pitfalls.
**Best LLM Code Tools (September 2026)** That's the English translation of the title. If you have the full article text you'd like translated as well, feel free to paste it and I'll translate it while preserving URLs, slugs, and markdown formatting.
Check out the ranking of the best LLMs for coding in September 2026: BenchLM scores, comparison and guide to choosing the right AI model.
iFlytek releases Spark X2.5: two open-source edge models with 1M context tokens — and a 293B arriving on September 7
iFlytek releases Spark X2.5: two open-source edge models with 1M context tokens. Also discover the 293B model planned for Sept 7.
GLM-5.3-Flash and Qwen3.8-Flash-Next: China releases two Flash models on the same day — the price war enters the speed race
GLM-5.3-Flash and Qwen3.8-Flash-Next: discover how China accelerates the AI model price war with these two simultaneous launches.
July 17: Gemini 3.5 Pro and Shanghai's WAIC collide — the day AI officially goes bipolar
On July 17, 2026, the Gemini 3.5 Pro launch and Shanghai WAIC illustrate two opposing visions. Discover this key day for AI.
GPT-Live : OpenAI launches full-duplex voice — AI agents can finally listen and speak at the same time
OpenAI launches GPT-Live with full-duplex voice. Discover how AI agents can finally listen and speak at the same time.
Meta Muse Spark 1.1 : Meta launches its first paid model and enters the agentic coding battle
Discover Meta Muse Spark 1.1, Meta's first paid model. The giant enters the agentic coding battle and changes strategy.
SpaceXAI's Grok 4.5: The 1,500-Parameter Billion Model Trained on Your Cursor Data Arrives to the Public
Discover Grok 4.5 from SpaceXAI, a 1,500B parameter AI model trained on Cursor data, launched to surpass OpenAI.
July 9, 2026: the most competitive day in AI history — three frontier labs, three public models at the same time
July 9, 2026 makes AI history: OpenAI, xAI, and Anthropic launch models simultaneously. Discover this AI race.
Best Free LLMs (July 2026)
Discover the top free LLMs ranking for July 2026. Gemini Flash, Llama 3.3: models with no credit card needed and stunning performance.
Fable 5 switches to credit-based billing — Anthropic's frontier model now costs $10 and $50 per million tokens
Anthropic changes Fable 5 access to credit billing. Discover the new frontier model pricing: $10 and $50 per million tokens.
Portugal launches Amália, its first sovereign open-source AI model — for 7 million euros
Discover Amália, Portugal's first sovereign open-source AI model, developed for just €7 million. A concrete European alternative.
ICML 2026 Seoul: 6,500+ papers accepted, ML enters the agentic era — key takeaways
Explore AI trends at ICML 2026 Seoul: over 6,500 accepted papers and the agentic era in machine learning.
Claude Sonnet 5: Anthropic's most agentic model, Opus performance at Sonnet price
OpenAI GPT-5.6: Sol, Terra et Luna — the model family that changes everything
Discover OpenAI GPT-5.6: Sol, Terra and Luna, the revolutionary model family under direct government control from June 26, 2026.
GPT-5.6 Sol: OpenAI launches the preview of a new model amid the early price war
Discover GPT-5.6 Sol, OpenAI's new preview shaking up the AI market amid a price war. Analysis and stakes of this launch.
Poolside Laguna M.1: the 225B open-source model for the coding agent, Apache 2.0
Discover Poolside Laguna M.1, a 225B-parameter open-source model under Apache 2.0, built to revolutionize coding agents.
FrontierCode: Cognition's benchmark that buries SWE-Bench and ranks code agents by the real quality of pull requests — Fable 5 at 46.3%, Opus 4.8 at 34.3%, GPT-5.5 at 25.5%
Discover FrontierCode, Cognition's new benchmark replacing SWE-Bench by evaluating the real quality of code agents' pull requests.
DeepSWE: the benchmark proving that code agents were cheating — Artificial Analysis buries SWE-Bench
Discover DeepSWE, the new benchmark replacing SWE-Bench, proving code agents were cheating. Analysis of the rankings upended by Artificial Anal
Gemini 3.5 Pro: countdown — 10 days before Google's deadline, 2 million tokens and Deep Think mode, the most anticipated model of the year (amidst a talent chaos)
Gemini 3.5 Pro: 10 days before Google's deadline, discover the rumors about its 2 million tokens and Deep Think mode amid a talent chaos.
GLM-5.2: The most powerful open weights model in the world — 753B MoE, 1M context, MIT license, the LLM landscape shifts
Discover GLM-5.2 from Z.ai: the world's most powerful open weights model. 753B MoE, 1M context & MIT license shaking up the LLM landscape.
CacheRL: A Qwen3-4B model achieves 92% accuracy in tool-calling with 100 times less compute than GPT-5
Discover CacheRL: a Qwen3-4B model hits 92% tool-calling accuracy with 100x less compute than GPT-5. AI revolution!
Best LLM Code (June 2026)
Discover the ultimate comparison of the best coding LLMs in June 2026. Analysis of agentic models capable of coding without human supervision.
Best Local LLMs (June 2026)
Discover the final ranking of the best local LLMs in June 2026. DeepSeek V4 Pro, Ollama: compare quality and privacy.
Kimi K2.7-Code : the 1T parameter open-source coding model that cuts 30% of reasoning tokens and beats Opus in tool use
Discover Kimi K2.7-Code, a 1T-parameter open-source coding model cutting reasoning tokens by 30% and outperforming Opus in tool use.
DeepSeek V4-Pro : the permanent 75% price drop accelerating the LLM war
DeepSeek V4-Pro permanently drops its price by 75%. Discover how this LLM model disrupts the market and accelerates the AI war.
Qwen3 Coder Next : the open-source model that runs on a 64 GB Mac and beats DeepSeek in coding
Discover Qwen3 Coder Next, the open-source model running on a 64GB Mac and beating DeepSeek at coding. A revolution for local code!
DiffusionGemma : Google releases the first open source diffusion text model — 4x faster than autoregressive
Discover DiffusionGemma: Google's first open-source diffusion text model, 4x faster than classic autoregressive approaches.
Best LLMs (June 2026)
Discover the full June 2026 best LLM ranking after the GPT-5.5 release. Compare autonomous AI models and their reasoning.
Claude Fable 5: Anthropic makes its Mythos model accessible to the public
Anthropic launches Claude Fable 5, the first public version of its Mythos model. Discover this model deemed too powerful and its explosive scores.
Best Free Llms (June 2026)
Discover the ranking of the best free LLMs in June 2026. Market analysis and comparison of uncensored AI models.
DeepSeek's DeepEP: the open source lib that optimizes GPU communication for large-scale MoE models
DeepSeek releases DeepEP, an open-source library that optimizes GPU communication to accelerate large-scale MoE model training.
NVIDIA Nemotron 3 Ultra 550B: The most powerful open-source model in the US arrives at Computex
Discover NVIDIA Nemotron 3 Ultra 550B, the most powerful US open-source model unveiled at Computex 2026 to rival China.
MiniMax M3: the Chinese open-weights model defying GPT-5.5 with 1M context and MSA architecture
Discover MiniMax M3, the Chinese open-weights model challenging GPT-5.5. It offers 1 million context tokens via MSA architecture.
DeepSeek V3.1: the silent revolution of open source arrives under the MIT license
DeepSeek V3.1 disrupts open source AI with a 671B parameter model under MIT license, with zero commercial restrictions.
Claude Opus 4.8: the model that dethrones GPT-5.5 — benchmarks, Dynamic Workflows, and the future of the coding agent
Anthropic's Claude Opus 4.8 dethrones GPT-5.5. Discover its benchmarks, the Dynamic Workflows system, and the coding agent revolution.
GPIC : Stanford releases 28 trillion pixels to train image generation models
Stanford releases GPIC, a 28-trillion-pixel dataset for training image generation models. Discover this permissive dataset.
LLMSurgeon: this ACL 2026 paper opens the black box of LLM pre-training
Discover LLMSurgeon, the ACL 2026 paper that opens the LLM pre-training black box to reveal their secret data mix.